Papers with long-horizon reasoning
Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Reinforcement learning (RL) has re-emerged as a natural approach for training interactive LLM agents in real-world environments. |
| Approach: | They propose a variant that operates on a turn-level MDP formulation, instead of the commonly used token-level one. |
| Outcome: | The proposed method is more robust than the widely used GRPO algorithm and more efficient than token-level MDPs. |
Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning (2026.acl-long)
Copied to clipboard
Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Jinhe Bi, Kristian Kersting, Jeff Z. Pan, Hinrich Schuetze, Volker Tresp, Yunpu Ma
| Challenge: | Large Language Models (LLMs) are stateless and limited by a finite context window, preventing them from maintaining knowledge across long conversations or evolving tasks. |
| Approach: | They propose a reinforcement learning framework that empowers LLMs to actively manage external memory through two specialized agents. |
| Outcome: | The proposed framework outperforms baselines and benchmarks across diverse question types, three benchmarks, and multiple model scales. |
Grounding Agent Memory in Contextual Intent (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models are deployed in long-horizon tasks that require agents to track interleaved goals, resolve references to prior information, and coordinate actions over extended trajectories. |
| Approach: | They propose an agentic memory system that indexes each trajectory step with a structured retrieval cue, contextual intent, and retrieves history by matching the current step’s intent. |
| Outcome: | The proposed system outperforms the strongest benchmark by 35.6%, with the largest gains as trajectory length increases. |
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning (2025.emnlp-main)
Copied to clipboard
Xiwen Liang, Min Lin, Weiqi Ruan, Rongtao Xu, Yuecheng Liu, Jiaqi Chen, Bingqian Lin, Yuzheng Zhuang, Xiaodan Liang
| Challenge: | Existing vision-language planning methods struggle with long-horizon reasoning in dynamic environments due to the difficulty of training models to generate high-quality reasoning processes. |
| Approach: | They propose a framework that enhances reasoning and action selection for long-horizon task planning through structured evaluation and optimized training. |
| Outcome: | The proposed framework outperforms existing methods on short-horizon tasks but struggles with long-horizon reasoning in dynamic environments. |
Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods handle long-term memory (LTM) and short-term (STM) as separate components, relying on heuristics or auxiliary controllers, which limits adaptability and end-to-end optimization. |
| Approach: | They propose a framework that integrates LTM and STM management directly into the agent's policy and propose 'agentic memory' to train such unified behaviors. |
| Outcome: | The proposed framework outperforms strong memory-augmented baselines on five long-horizon benchmarks and achieves higher-quality long-term memory and more efficient context usage. |
Memory Matters More: Event-Centric Memory as a Logic Map for Agent Searching and Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for storing and retrieving memory are limited by shallow semantic retrieval. |
| Approach: | They propose a memory mechanism that organizes and retrieves past experiences to support decision-making. |
| Outcome: | Experiments on LoCoMo and NarrativeQA show that CompassMem improves retrieval and reasoning performance across multiple backbone models. |
TOOLCAD: Exploring Tool-Using Large Language Models in Text-to-CAD Generation with Reinforcement Learning (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown remarkable advances in enabling language agents to tackle real-world tasks. |
| Approach: | They propose a tool-using agent-based CAD framework that automates text-to-CAD modeling . they propose an interactive CAD gym to roll out reasoning and tool-augmented interaction trajectories with the CAD engine . |
| Outcome: | The proposed framework can generalize across complex modeling tasks, supporting their open-source counterparts. |
CIRAG: Construction–Integration Retrieval and Adaptive Generation for Multi-hop Question Answering (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for iterative retrieval-augmented generation (iRAG) suffer from greedy single-path expansion and granularity–demand mismatch . |
| Approach: | They propose a model that constructs candidate triples and history-conditionally integrates them to distill core triples to generate the next-hop query. |
| Outcome: | The proposed model mitigates the greedy single-path expansion and granularity–demand mismatch by preserving multiple plausible evidence chains. |
Octopus: Gated Selective Attention for Memory-Bounded Long-Context Inference in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Subquadratic architectures rely on aggressive state compression that degrades performance on complex reasoning tasks. |
| Approach: | They propose a framework that confers fixed-memory inference onto pretrained Transformers . they use a learnable module that enforces an adaptive sparsity policy over the context history . |
| Outcome: | The proposed framework outperforms state-of-the-art linearized baselines on the GSM8K benchmark by over 36 points under identical memory constraints. |
CheckRLM: Effective Knowledge–Thought Coherence Checking in Retrieval-Augmented Reasoning (2026.acl-long)
Copied to clipboard
Dingling Xu, Ruobing Wang, Qingfei Zhao, Yukun Yan, Zhichun Wang, Daren Zha, Shi Yu, Zhenghao Liu, Shuo Wang, Xu Han, Maosong Sun
| Challenge: | Reasoning Language Models (RLMs) have improved performance on complex tasks by extending the reasoning chain, but they are prone to factual errors, especially in knowledge-intensive tasks. |
| Approach: | They propose a framework that improves the reliability of the reasoning process by timely checking and correcting factual errors. |
| Outcome: | The proposed framework outperforms baselines and shows that it mitigates error accumulation with lower costs. |
BMAM: Brain-inspired Multi-Agent Memory Framework (2026.findings-acl)
Copied to clipboard
| Challenge: | Language-model-based agents operating over extended interactions face persistent challenges in preserving temporally grounded information and maintaining behavioral consistency across sessions. |
| Approach: | They propose a general-purpose memory architecture that decomposes agent memory into six components that operate at complementary time scales. |
| Outcome: | BMAM outperforms memory-augmented baselines on LoCoMo benchmarks with 78.45% accuracy . a targeted refinement of the temporal-trigger heuristics raises LongMemEval multi-session accuracy from 45.2% to 56.4%. |